Papers by Pedro Ortiz Suarez

8 papers
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing web crawling pipelines are used to collect large corpora raw data, but the main way to collect such data is through manual data extraction.
Approach: They propose to use a web crawler to extract and classify data from a multilingual web corpus and an automated annotation pipeline to improve it.
Outcome: The proposed version of OSCAR could be used to pre-train large generative language models and other applications in Natural Language Processing and Digital Humanities.
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)

Copied to clipboard

Challenge: Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available.
Approach: They propose to use a contextualised language model to analyse historical states of language in French.
Outcome: The proposed model is based on a corpus of historical texts and is evaluated with an NLP task.
A Data-driven Approach to Named Entity Recognition for Early Modern French (2022.coling-1)

Copied to clipboard

Challenge: Named entity recognition is an important task in natural language processing.
Approach: They propose to use a data-driven approach to identify historical French with fine-grained annotations instead of a specialised architecture to tackle particularities.
Outcome: The proposed corpus is larger than the most popular NER evaluation corpora for both Contemporary English and French.
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies show that multimodal large language models can learn from text-image data.
Approach: They propose to train multimodal large language models on large amounts of text-image data . they also show a boost in few-shot learning performance across various multilingual tasks .
Outcome: The proposed dataset is not public and is only in English . it is the first large-scale multilingual and multimodal document corpus crawled from the web.
BERTrade: Using Contextual Embeddings to Parse Old French (2022.lrec-1)

Copied to clipboard

Challenge: a growing interest in digital humanities for automatic processing and annotation of historical texts is generating new models for historical languages.
Approach: They use POS-tagging and dependency parsing to evaluate contextual word embedding models . Old French is one of the historical languages for which they have the largest amount of syntactically annotated data .
Outcome: The proposed model can be used to improve performance in Old French, the authors show . they use POS-tagging and dependency parsing to evaluate the model's quality .
Tokenizer Choice For LLM Training: Negligible or Crucial? (2024.findings-naacl)

Copied to clipboard

Challenge: Recent success of large language models has been driven by curating the training dataset composition, scaling of model architectures and advancements in pretraining objectives, leaving tokenizer influence as a blind spot.
Approach: They conduct a comprehensive study on the influence of tokenizer choice on LLM downstream performance by training 24 mono- and multilingual LLMs at a 2.6B parameter scale.
Outcome: The proposed model can significantly impact the model's downstream performance and training costs.
Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem (2025.findings-emnlp)

Copied to clipboard

Challenge: a new pipeline can be used to create corpora for over-looked languages .
Approach: We propose a new pipeline that can filter a single snapshot in twohours.
Outcome: The proposed pipeline can filter a single snapshot in twohours.
A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages (2024.lrec-main)

Copied to clipboard

Challenge: CATalog 1.0 is the largest text corpus in Catalan to date . CURATE is a pipeline that can be parallelizable to run in high performance clusters .
Approach: They propose a data pipeline that uses binary filters to filter documents based on text quality . they optimised the pipeline to run in high performance clusters .
Outcome: The proposed pipeline is optimized for high performance cluster environments and runs in high performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations